Papers with bilingual language models
Pula: Training Large Language Models for Setswana (2025.naacl-long)
Copied to clipboard
| Challenge: | Setswana is a Bantu language spoken by an estimated five to ten million people worldwide. |
| Approach: | They propose to make setswana-based models available for the first time using data available from setswa and setswegian databases. |
| Outcome: | The proposed models outperform GPT-4o and Gemini 1.5 Pro on English-Setswana translation tasks and achieve state-of-the-art performance on Setswanan reasoning tasks. |
How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders (2025.findings-emnlp)
Copied to clipboard
Tatsuro Inaba, Go Kamoda, Kentaro Inui, Masaru Isonuma, Yusuke Miyao, Yohei Oseki, Yu Takagi, Benjamin Heinzerling
| Challenge: | Using sparse autoencoders, we explore how bilingual language models develop complex internal representations. |
| Approach: | They employ sparse autoencoders to analyze bilingual language models' internal representations. |
| Outcome: | The proposed method integrates decomposed representations from a fully trained model into a mid-training model. |